Back

Biology Methods and Protocols

Oxford University Press (OUP)

Preprints posted in the last 30 days, ranked by how well they match Biology Methods and Protocols's content profile, based on 61 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.

1
Automated Detection of Extrahepatic Bile Duct Stones on Intraoperative Cholangiography Using Deep Learning

He, Y.; Bloom, M.; Mirshojae, S.; Noel, L.; Qureshi, T.; Xie, Y.; Phillips, E.; Li, D.; Huang, X.

2026-08-24 surgery 10.64898/2026.08.20.26360965 medRxiv
Top 0.1%
6.7%
Show abstract

Objective: To evaluate the case-level performance of deep-learning segmentation models for detecting extrahepatic bile duct stones on representative intraoperative cholangiography images (IOC) and to characterize the completeness of individual-stone localization. Background: Retained bile duct stones can cause biliary obstruction, cholangitis, and pancreatitis. However false-positive interpretation of filling defects may prompt additional downstream procedures. Computer vision has been applied to biliary anatomy recognition and IOC adequacy assessment, but patient-level stone detection and individual-stone localization remain insufficiently studyed. Methods: Representative IOC images were annotated for extrahepatic biliary anatomy and stones, with case-level stone status established using a composite clinical reference standard. Two deep-learning models were developed to delineate the common bile duct and common hepatic duct and to detect and localize stones. Case-level diagnostic performance was evaluated against the composite clinical reference standard, and individual-stone localization was evaluated against expert-reviewed annotations. Results: On the held-out 125 patients test set, MiT-B2-UNet identified 23 of 25 stone-positive cases and 95 of 100 stone-negative cases, corresponding to a sensitivity of 0.920, specificity of 0.950, and AUC of 0.986. nnU-Net identified 19 of 25 stone-positive cases and 98 of 100 stone-negative cases, corresponding to a sensitivity of 0.760, specificity of 0.980, and AUC of 0.959. At the individual-stone level, MiT-B2-UNet and nnU-Net localized 31 of 59 and 25 of 59 annotated stones, respectively; all annotated stones were localized in 13 of 25 and 12 of 25 stone-positive cases. Conclusions: Deep-learning models can identify stone-positive IOC cases and localize individual stones. This technology may help inte

2
A Deep Learning-Derived Insulin Resistance Index for Cardiovascular Risk Prediction: A Prospective Cohort Study with External Validation in Chinese and US Populations

Mao, Y.; Lin, J.; Zhou, A.; Zeng, S.; Yang, D.; Lin, W.; Wen, J.; Yang, W.; Chen, G.

2026-08-12 endocrinology 10.64898/2026.08.10.26360145 medRxiv
Top 0.1%
5.4%
Show abstract

Background Existing insulin resistance (IR) indices are predominantly developed in diabetic cohorts, limiting their generalizability. We developed a novel deep neural network-derived IR index (DNN-IR) using a Mixture-of-Experts (MoE) framework and evaluated its predictive performance for incident cardiovascular disease (CVD) and mortality in general populations. Methods We utilized data from three cohorts: the cross-sectional REACTION study (Fujian subcohort, 2011-2012) for DNN-IR derivation and internal validation; and two prospective cohorts, NHANES (1999-2018, linked to the National Death Index) and CHARLS (2011-2018), for external validation. The DNN-IR was developed using a deep learning model based on a Mixture-of-Experts (MoE) architecture, trained on the REACTION dataset. We evaluated the DNN-IR's utility in predicting incident CVD, cardiovascular mortality, and non-cardiovascular mortality among 13,889 NHANES and 7,047 CHARLS participants. Predictive performance was assessed via the area under the receiver operating characteristic curve (AUC). Multivariable logistic regression, restricted cubic splines, and Kaplan-Meier analyses characterized the associations between DNN-IR and clinical outcomes. Results In the REACTION cohort, DNN-IR demonstrated superior predictive performance for atherosclerotic outcomes, achieving AUROCs of 0.89 (training) and 0.84 (internal validation). In the external CHARLS cohort (median follow-up: 7 years; 1,135 incident CVD cases [16.1%]), DNN-IR yielded AUROCs of 0.72 for incident CVD and 0.77 for all-cause mortality. Fully adjusted models showed that each 1-SD increment in DNN-IR was associated with a 23% higher CVD risk (OR=1.23, 95% CI: 1.14-1.32), exhibiting a predominantly linear dose-response relationship (P-nonlinearity=0.453). In NHANES, DNN-IR robustly predicted cardiovascular (AUROC=0.77) and all-cause mortality (AUROC=0.72), alongside specific mortalities like diabetes (0.91), Alzheimer's disease (0.88), and kidney disease (0.96). Higher DNN-IR levels correlated with stepwise increases in cumulative mortality (log-rank P<0.001). Conclusions The MoE-derived DNN-IR index demonstrated robust and stable performance in predicting atherosclerosis, incident CVD, cardiovascular mortality, and all-cause mortality in the general population. Further validation in larger, more diverse cohorts is warranted to support its broad clinical applicability.

3
High thoughput fluorometric nucleic acid quantification using qPCR instruments

Meerson, A.

2026-08-06 molecular biology 10.64898/2026.08.01.742208 medRxiv
Top 0.1%
5.0%
Show abstract

To explore adapting qPCR systems for end-point nucleic acid quantification using dyes such as SYTO-9, we quantified serial dilutions of DNA and RNA standards in the range of 0.75 - 200 ng/{micro}l on 384-well qPCR devices. SYTO-9 fluorescence was successfully measured using standard SYBR Green settings. Blank-subtracted relative SYTO-9 signal showed a logarithmic dependence on DNA/RNA concentration (R2 > 0.95). Measurements were highly stable with different incubation times, temperatures of up to 95{degrees}C, and photobleaching. The described approach is a valuable QC option for high-throughput DNA/RNA isolations and could be adapted to additional fluorometric assays beyond nucleic acids.

4
Deep Learning of Fluorescence Lifetime Imaging Ophthalmoscopy for Type 2 Diabetes Classification

Kwon, S.; Lee, C. S.; Lee, A. Y.; Zhang, L.

2026-08-06 endocrinology 10.64898/2026.08.04.26359728 medRxiv
Top 0.1%
4.9%
Show abstract

Purpose: To evaluate whether fluorescence lifetime imaging ophthalmoscopy (FLIO) combined with deep learning can detect metabolic signatures for classification of type 2 diabetes mellitus (T2DM). Design: Cross-sectional analysis of participants included AI-READI dataset (version 3) with FLIO imaging and and hemoglobin A1c (HbA1c) measurement. Subjects: 1,783 participants from the AI-READI dataset (version 3) with HbA1c measurements and FLIO imaging scans (6,912 total): 671 normoglycemic, 726 prediabetic, and 386 diabetic. Methods: Mean fluorescence lifetime maps were generated using a center-of-mass approach and used as inputs to AI models. We trained convolutional neural networks (CNNs), ResNet-18, and XGBoost under three-class (normal, prediabetic, diabetic) and two binary (normal vs. impaired; normal vs. diabetic) classification schemes, using nested 5-fold cross-validation with participant-level grouping. Main Outcome Measures: Macro-averaged accuracy, F1 score, area under the receiver operating characteristic curve (AUROC), sensitivity, specificity, and positive predictive value (PPV). Results: Group-averaged lifetime maps demonstrated consistent spatial differences across glycemic groups, with progressively longer lifetimes from normal to diabetic participants. The CNN achieved the best overall performance in the 3-class classification (accuracy 0.41 +/- 0.03, F1 score 0.39 +/- 0.02, AUROC 0.58 +/- 0.02), compared to the random classifier for 3-class classification (AUROC = 0.50; accuracy = F1 = 0.33). ResNet-18 and XGBoost showed similar performance (AUROC 0.53-0.58). Confusion matrices revealed substantial overlap between classes, with frequent misclassification toward the prediabetes group. Binary reformulation (normal vs. diabetic) improved performance substantially, with the CNN resulting in AUROC 0.63 +/- 0.02 and XGBoost 0.67 +/- 0.07. Conclusions: FLIO-derived lifetime maps capture metabolic signals associated with glycemic status but yield modest classification performance with current AI models. These findings highlight both the potential and the challenges of using FLIO for early metabolic screening and monitoring, informing future development of clinically applicable imaging biomarkers.

5
Does Data Preprocessing Affect Tree-Based Super Learners? An Investigation of Ensemble Optimization and Oracle Properties in Clinical Classification.

Darko, R.; Dwumah, D.; Agyapong, K. S.; Agyenim-Boateng, Y.; Darko Anim, R.; Wisdom Jakper, J.; Owusu-Ansah, N. K.; Owusu-Ansah, R.

2026-08-24 health informatics 10.64898/2026.08.20.26360880 medRxiv
Top 0.1%
4.8%
Show abstract

Machine learning workflows frequently incorporate data preprocessing to enhance predictive performance. However, the need for Super Learner ensembles made up only of preprocessing-invariant tree-based algorithms remains unexplored. Using three benchmark clinical classification datasets, this study examined how preprocessing affected the Super Learner's prediction performance, learner weight distribution, and oracle behavior. The Heart Disease (207 observations), Indian Liver Patient Dataset (583 observations), and Pima Indians Diabetes (768 observations) datasets were used to create a Super Learner ensemble model that included Classification and Regression Trees (CART), Random Forest, Ranger, and Extreme Gradient Boosting (XGBoost). Models were evaluated under raw and preprocessed data conditions using repeated cross-validation. Predictive performance was assessed using the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), and Brier score. Learner weight allocation and Oracle Gap were compared using paired Wilcoxon signed-rank tests with Benjamini-Hochberg adjustment. Preprocessing produced negligible changes in predictive performance for the Heart Disease and Pima datasets. For the ILPD dataset, preprocessing significantly improved AUC (0.746 to 0.752; adjusted p = 0.0017) and reduced the Brier score (0.177 to 0.175; adjusted p < 0.001). Learner weights remained largely stable, although Random Forest replaced Ranger as the dominant learner for the Heart Disease dataset. Oracle Gaps remained extremely small (<0.002) across all datasets and did not differ significantly between preprocessing conditions. Preprocessing provides limited benefit for Super Learner ensembles composed of preprocessing-invariant learners and does not materially alter their oracle behavior. Preprocessing decisions should therefore be guided by dataset characteristics rather than adopted as a universal modelling practice.

6
Multimodal artificial intelligence for personalized hepatocellular carcinoma treatment strategy selection

Feng, W.; Liu, S.; Yang, Z.; Tao, Y.; Gu, X.; Jin, W.

2026-08-25 health informatics 10.64898/2026.08.21.26361067 medRxiv
Top 0.2%
4.3%
Show abstract

Background Hepatocellular carcinoma (HCC) treatment selection demands nuanced integration of heterogeneous patient data, yet prevailing predictive models rely on restricted data modalities and oversimplified therapeutic frameworks, compromising clinical translation. Objective We developed and validated a multimodal artificial intelligence framework to guide optimal treatment strategy selection across the full spectrum of HCC interventions. Methods This retrospective study comprised 1,043 HCC patients (development cohort, January 2017-December 2023) and 55 external validation patients (2023) from Wuxi Peoples Hospital. We engineered Embedding-Augmented Extra Trees (ET-Emb), a novel model fusing structured clinical variables with contextual text embeddings derived from medical histories and radiology reports. ET-Emb quantifies probabilities for five primary treatments: open/laparoscopic resection, transarterial chemoembolization, radiofrequency ablation (RFA), and chemotherapy. Model performance was rigorously assessed via 10-fold cross-validation and external validation using ROC-AUC and PR-AUC metrics. Results ET-Emb demonstrated robust performance in the development cohort (ROC-AUC: 0.84 {+/-} 0.04; PR-AUC: 0.55 {+/-} 0.06), significantly outperforming established benchmarks. This generalizability was preserved in external validation (ROC-AUC: 0.77 {+/-} 0.02; PR-AUC: 0.47 {+/-} 0.03). SHAP analysis identified textual clinical narratives and socioeconomic determinants as critical predictive drivers. Conclusions By unifying structured and unstructured data modalities, ET-Emb delivers accurate, multi-treatment strategy prediction for HCC. Its clinical validity and the demonstrated significance of textual features establish multimodal AI as an essential paradigm for simulating complex oncological decision-making, positioning ET-Emb as a transformative tool for precision HCC management.

7
A Cesium Chloride Gradient Ultracentrifugation-Based Method for the Isolation of DNA from Diverse Recalcitrant Plant Species for Nanopore Sequencing

Labbancz, J.; Dhingra, A.

2026-08-21 molecular biology 10.64898/2026.08.18.745475 medRxiv
Top 0.2%
4.3%
Show abstract

Developments in Nanopore sequencing have enabled telomere to telomere genomic assembly as a routine technique in genomic research. Nanopore DNA sequencing for genomic assembly is typically performed on native DNA molecules, making it particularly sensitive to the quality of input DNA, with contaminating molecules limiting data yields and reducing read quality. As pangenome analysis gains interest, particularly in non-model plant species which are often rich in inhibitory secondary metabolites, the development of methods which can improve the quality and throughput of nanopore sequencing is essential. Here we describe a method for isolation of total DNA from the leaf tissues of diverse Viridiplantae species. The initial lysis buffer consists of a modified CTAB buffer, incorporating dimethyl sulfoxide for the reduction of viscosity, which can be problematic in many plant DNA preparations. An organic extraction with 2-butoxyethanol is utilized to further extract phenolic compounds which may be sufficiently hydrophilic to evade chloroform extraction, while reducing aqueous phase volume. Further cleanup via cesium chloride (CsCl) ultracentrifugation is performed to minimize the carryover of residual contaminating macromolecules. Samples prepared using this method are of consistent high quality, even when extracted from challenging late season leaf tissue or secondary metabolite rich species. Sequencing results from samples prepared by this method outperform those obtained from typical modified CTAB DNA isolation techniques in both quantity and quality. We tested sequencing performance from Vitis DNA isolated using a modified CTAB method and Vitis DNA isolated using the CsCl ultracentrifugation-based method described here. DNA isolated via the method described here produced 83% more >Q10 sequence data (52.61 Gb vs. 28.8 Gb), resulted in a 60% greater read N50 despite more handling steps (32.78kb vs. 20.45kb), and resulted in a higher modal read quality (Q27 vs. Q24). The consistency of this method across diverse plant taxa suggests its use as a general method for DNA isolation prior to Nanopore sequencing and genomic assembly for diverse plant taxa.

8
Modular Robotic NOSES-II for Mid-Rectal Cancer: A Preliminary Feasibility Study of a Structured Surgical Procedure

Liu, P.; Zhang, l.; Yu, K.; Lu, X.; Li, W.

2026-08-17 surgery 10.64898/2026.08.15.26359890 medRxiv
Top 0.2%
3.8%
Show abstract

Background and Objectives: Robotic-assisted natural orifice specimen extraction surgery (NOSES) is a minimally invasive approach for mid-rectal cancer, but it is technically complex and lacks a standardized operational framework. This study aimed to propose and preliminarily validate a structured modular robotic NOSES II surgical procedure for mid-rectal cancer. Methods: This was a retrospective observational study that consecutively enrolled 11 patients with mid-rectal cancer who underwent modular robotic NOSES-II surgery at our center between December 2024 and August 2025. All procedures were performed by the same surgical team strictly following the predefined six-module structured surgical protocol. Perioperative indicators, pathological outcomes, and patient-reported outcomes (PROs) at 4 to 8 months postoperatively were collected. Key techniques were analyzed via high-definition surgical videos. Results: All 11 procedures were completed successfully without conversion to open surgery. The mean operative time was 299.6 plus or minus 54.2 minutes, and the mean intraoperative blood loss was 94.1 plus or minus50.6 mL. The mean number of harvested lymph nodes was 17.4 plus or minus 8.4, with a 100% R0 resection rate. No severe complications (Clavien-Dindo grade [&ge;] III) occurred. The mean postoperative hospital stay was 8.2 plus or minus 1.5 days. Postoperative PROs indicated good defecatory function, urinary function, and overall quality of life. Conclusions: Preliminary findings demonstrate that the structured modular robotic NOSES-II procedure is safe and feasible for the treatment of mid-rectal cancer. This modular protocol provides a clear and reproducible technical framework for this complex procedure, and it is expected to achieve favorable functional preservation while ensuring oncological radicality.

9
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.2%
3.5%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

10
Cohort profile: KaroLiver, a population-based cohort of patients receiving curative liver-directed treatment for colorectal liver metastases at a Swedish tertiary centre

Gerling, M.; Moro, C. F.; Limbecker, C.; Viljamaa, A.; Harrizi, S.; Hamidi, Y.; Hailer, A.-K.; Sterner, J.; Sparrelid, E.; Bozoky, L.; Tidholm Qvist, E.; Baumgartner, R.; Salmonson Schaad, M.; Bozoky, B.; Geyer, N.; Engstrand, J.

2026-08-27 surgery 10.64898/2026.08.24.26361235 medRxiv
Top 0.2%
3.5%
Show abstract

Purpose The Karolinska Liver Metastases (KaroLiver) cohort was established to investigate associations between clinical characteristics and histopathological features in patients treated with curative intent for colorectal cancer liver metastases (CRLM). The cohort combines whole-slide digital histopathology images with detailed oncological, surgical, radiological and survival data, enabling comprehensive analyses of treatment trajectories, clinical outcomes and metastatic tumour biology. Participants KaroLiver is a retrospective observational cohort comprising all consecutive patients who received curative-intent, liver-directed treatment for CRLM at Karolinska University Hospital in Stockholm, Sweden. The hospital is the primary regional referral centre for HPB surgery, serving the population of approximately 2.5 million people in the Stockholm-Gotland healthcare region. Patient enrolment is continuously updated in accordance with amended ethical approvals and evolving scientific questions. The cohort currently comprises 811 patients who underwent 1204 liver interventions between February 2012 and January 2022. Detailed clinical, oncological, surgical, pathological, molecular, recurrence and survival data are collected. Findings to date Median overall survival (OS) in the current cohort is 51.0 months (95% CI 46.2-57.1 months), and median recurrence-free survival (RFS) is 10.4 months (95% CI 9.2-11.9 months). The five-year OS rate is 44.9% (95% CI 41.2-48.8%). Studies using the cohort have so far identified a liver injury-derived stromal capsule in a subset of metastases, associated with improved survival. The cohort has also enabled the identification of histopathological markers of tumour biology, sex-based differences in treatment and survival, and associations between post-hepatectomy liver failure and oncological outcomes. Future plans Current research priorities include advanced histology-based prognostic scoring, sex differences in recurrence and retreatment, tumour biology and outcomes in early-onset versus average-onset CRLM, as well as CT- and MRI-based radiomics, all integrated within KaroLiver's histopathological framework. Data sharing is supported, given that regulatory requirements are met. Retrospective accrual and outcome updates will continue for current and future studies, subject to the required approvals.

11
Multimodal Large Language Models vs. Medical Doctors in Degenerative Lumbar Spine Surgery: A Retrospective Decision Concordance Study of 147 Patients

Hamdan, M.; Harati, A.; Al-Bakheet, A.; Fuetterer, I.; Alshaer, I.

2026-08-06 surgery 10.64898/2026.08.04.26359718 medRxiv
Top 0.3%
3.3%
Show abstract

Objective: To evaluate decision concordance between commercially available multimodal large language models (LLMs), resident doctors, and senior-surgeon ground truth for surgical indication and spinal level in degenerative lumbar spine disease. Methods: We retrospectively analyzed 147 consecutive patients. Each case included clinical documentation and MRI presented as two composite PNG images. Two resident doctors and three multimodal LLMs (GPT 5.5, Claude Sonnet 4.6, Gemini 3.1 Pro) independently assessed operative versus conservative management and, if operative, the surgical level. Analyses used Cochran's Q, McNemar tests with Holm correction, and Bayesian methods. Results: LLMs achieved higher therapy-decision accuracy (66.0%-68.0%; 97-100/147) than residents (54.4%; 80/147) but over-recommended surgery. Conditional level accuracy when surgery was correctly indicated was 71.4% (20/28) for residents versus 33.3%-41.1% for LLMs. Conclusion: Off-the-shelf multimodal LLMs approximate human performance for binary surgical indication but remain inferior for precise level localization. These results establish a practice-relevant baseline of spatial reasoning limitations for tools already used by patients and junior doctors.

12
A software package for simple and rigorous survival machine learning analysis in biomedical research

Pybus, A.; Qiu, J.; Morais Lyra, P. C.; Dang, K.; Narvaez-Bandera, I.; Jolaogun, T.; Goecks, J.

2026-08-10 health informatics 10.64898/2026.08.05.26359034 medRxiv
Top 0.3%
3.2%
Show abstract

Survival analysis is a fundamental technique in biomedical research for modeling time-to-event data. It enables the identification of prognostic factors in disease, compares survival outcomes across treatment groups, and performs targeted treatment selection. A variety of machine learning (ML) approaches to survival analysis have emerged to complement classical statistical methods, especially for high-dimensional datasets with complex, nonlinear interactions between features. However, using survival ML methods requires addressing challenges such as censoring-unaware evaluation, overfitting, selecting performance metrics, and data leakage. To address these and other difficulties in using survival ML models, we developed the mlsurv software package. mlsurv is an open-source Python package built around three major design principles: 1) methodological rigor, including evidence-based model selection, leakage-free pipelines, and multi-metric evaluation, 2) multi-scale evaluation and interpretation, including population and subpopulation evaluation, patient-level explanations, and feature analysis, and 3) automated trust and transparency, including limitation flagging and TRIPOD+AI-aligned reporting. mlsurv bundles ten models spanning linear, ensemble, kernel, and deep learning families within a unified software package. We demonstrate mlsurv on the Chowell immunotherapy cohort (n=1,479). The survival-trained models achieve a test concordance index of 0.73 for overall survival prediction. Further, risk scores strongly correlate with the response-trained LORIS clinical score (|{rho}| up to 0.84), reflecting the overlap between prognostic and predictive signal. mlsurv enables biomedical researchers to conduct rigorous, multi-model survival analysis and benchmarking using minimal code with default best practices rather than implementing custom scripts and methodological safeguards from scratch.

13
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 0.3%
2.8%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.

14
Cross-Cohort Evaluation of NanoString nCounter Data for Recurrence Prediction in Colorectal Cancer

Quarles Van Ufford, P.; Bojesen, R. D.; Olsen, L. R.; Gogenur, I.; Lund, O.

2026-08-17 oncology 10.64898/2026.08.13.26360359 medRxiv
Top 0.4%
2.6%
Show abstract

Gene expression-based prognostic models have shown promise for predicting recurrence in colorectal cancer (CRC), but their clinical implementation remains limited. The NanoString nCounter platform provides a practical alternative to RNA sequencing and microarrays through standardized, cost-effective gene expression profiling that is compatible with routine clinical samples. In this study, we evaluated whether NanoString nCounter gene expression data improve prediction of recurrence following curative CRC surgery. Gene expression profiles from the NanoString PanCancer IO 360 panel were analyzed in two independent CRC cohorts (cohort A, n = 189; cohort B, n = 131). Differential gene expression analyses and Cox proportional hazards models were used to assess the prognostic value of gene expression alone and in combination with established clinical risk factors. Model performance was evaluated by five-fold cross-validation and external validation between cohorts using the concordance index (C-index) and Kaplan-Meier risk stratification. The two cohorts differed significantly in recurrence-free survival, and differential expression analysis demonstrated marked cohort-specific transcriptional patterns. Ninety-one recurrence-associated genes were identified in cohort A, whereas no significant genes were detected in cohort B, with poor agreement in gene-level differential expression between cohorts (Pearson r = 0.128). Across all prediction models, external performance was modest, and inclusion of gene expression data did not improve prediction beyond clinical variables. The clinical baseline model, incorporating age, UICC stage, and tumor site, consistently achieved the highest cross-cohort performance, with UICC stage emerging as the strongest predictor of recurrence. Although overall discrimination was moderate, the baseline model successfully stratified patients into significantly different high- and low-risk groups across cohorts. These findings indicate that prognostic gene expression signatures derived from NanoString data showed limited reproducibility across independent cohorts and provided little additional predictive value beyond established clinical factors. The results highlight the importance of external validation and suggest that robust clinical variables remain the most reliable predictors of recurrence risk in this setting.

15
Algorithmic Ascertainment of Cause of Death from Longitudinal Real-World Medical Claims Data: Development and Validation

McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360606 medRxiv
Top 0.4%
2.4%
Show abstract

This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.

16
Local retraining mitigates domain shift in sepsis prediction: Lessons from translating a neonatal model to mixed intensive care data

Champeaux, S. A.; Booth, J.; Brown, A.; Sebire, N. J.; Drobnjak, I.; Bowyer, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360666 medRxiv
Top 0.5%
2.4%
Show abstract

Background: Machine learning models leveraging electronic health records (EHRs) can support earlier detection of sepsis in intensive care units (ICUs). However, their clinical utility depends on reproducibility across institutions and patient populations. Building on a published pipeline from the Children's Hospital of Philadelphia (CHOP), this study examines how a neonatal sepsis prediction framework performs and can be adapted to a range of intensive care environments, paediatric, cardiac, and neonatal, at Great Ormond Street Hospital (GOSH). Methods: We extracted de-identified ICU EHR data from GOSH and applied feature derivation, unit harmonisation, and temporal sampling to align with the CHOP dataset used by Masino et al. (2019). Seven classifiers were first evaluated using CHOP-trained weights to characterise cross-domain behaviour and then retrained on local data to assess recoverability and site-specific adaptation. Model discrimination was summarised by AUC and F1, and learning curves were used to explore sample efficiency and bias-variance dynamics. Results: Models achieved strong discrimination on the CHOP neonatal cohort but demonstrated reduced performance when transferred to the mixed GOSH ICU population, reflecting anticipated domain and population shift. Retraining on GOSH data restored discrimination (AUC range 0.69-0.86), with Gradient Boosting (AUC 0.86 vs AUC 0.87 at CHOP) and KNN (AUC 0.80 vs AUC 0.79 at CHOP) models performing comparably to their CHOP benchmarks. DeLong's test confirmed statistically significant gains across all classifiers (p < 0.001). Conclusion: ICU cohort and baseline demographic differences between CHOP and GOSH introduced domain shift that limited direct model transfer. Elements of the original preprocessing pipeline could not be reproduced, further constraining transportability. Yet, retraining on local data restored high discrimination, showing that the modelling framework remains robust when re-estimated in new settings. These results highlight local adaptation as a practical route to recover performance and support safe, generalisable deployment of clinical prediction models in mixed clinical environments.

17
Video-based gait analysis using pose estimation can quantify gait differences among non-frail, pre-frail, and frail older adults

Burch, K.; Hamkins, J.; McDaniel, L.; Castro e Costa, A. R.; Yang, Z.; Stenum, J.; Pagliocchini, A.; Szczesny, C.; Langdon, J.; Chellappa, R.; Abadir, P.; Roemmich, R.

2026-08-07 geriatric medicine 10.64898/2026.08.04.26359742 medRxiv
Top 0.5%
2.3%
Show abstract

Frailty is a common consequence of aging that makes individuals increasingly susceptible to adverse health outcomes. Frailty screening can identify pre-frail and frail individuals to prescribe interventions or inform clinical decision making to prevent or slow additional frailty progression. Objective, scalable, and automated frailty assessments may expedite and improve clinical frailty screening. Here, we leveraged human pose estimation for video-based gait analysis in older adults who were non-frail, pre-frail, and frail. We focused on gait because slow walking speed is key diagnostic criteria of frailty, and many gait deviations are often observed in older adults with frailty. We collected videos of 68 older adults (25 non-frail, 25 pre-frail, 18 frail) walking at both self-selected and fast paces and used an established pose estimation-based gait analysis approach to measure and compare gait parameters across frailty statuses. Pose estimation-based step time measurements were strongly correlated with manual annotations (self-selected: R2=0.93, fast: R2=0.80) and showed tight Bland-Altman limits of agreement (self-selected: -0.082 to 0.052s, fast: -0.114 to 0.110s), establishing validity of this video-based gait analysis approach in older adults. We then identified a series of cross-sectional differences in spatiotemporal gait parameters among non-frail, pre-frail, and frail older adults, demonstrating that video-based gait analysis can be useful for measuring gait differences across frailty statuses. This study demonstrates the potential of video-based pose estimation for scalable gait tracking across frailty statuses in older adults.

18
Large language model-augmented implicit surgical video review

Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.

2026-08-31 surgery 10.64898/2026.08.25.26361071 medRxiv
Top 0.5%
2.3%
Show abstract

Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.

19
Predicting Subjective Cognitive Decline on Future BRFSS Survey Years: An Open Multi-Language Machine Learning Benchmark

Nguyen, T. T.; Nguyen, T. D.

2026-08-21 public and global health 10.64898/2026.08.18.26360709 medRxiv
Top 0.5%
2.2%
Show abstract

Background and Objectives: Subjective cognitive decline (SCD), self-reported worsening confusion or memory over the past year, is a common early marker of cognitive concern with relevance for Alzheimer's disease prevention and population health. Population-based machine learning benchmarks that respect temporal drift in public health surveillance remain limited. We developed a reusable multi-language prediction and interpretability framework for SCD using Behavioral Risk Factor Surveillance System (BRFSS) Cognitive Decline data. Methods: We analyzed pooled national (n = 298,944) and New York (n = 30,366) cohorts with chronological train (2015-2019), validation (national: 2020-2022; New York: 2020-2021), and locked test (2023-2024) splits. Nested LASSO identified stable predictors. Sixteen machine learning algorithms were compared under year-grouped cross-validation with SMOTE restricted to training folds. Four end-to-end Python/R pipelines (single-model or soft-voting) used validation-only isotonic calibration and Youden thresholding. Primary reporting pipelines were prespecified before test unlock (national: R tidymodels single-model; New York: Python single-model); algorithms within each pipeline were chosen by validation ROC-AUC. Post-hoc GLMs (national unweighted; New York design-weighted) and two training-only knowledge-graph layers supported interpretability. Results: Locked-test discrimination was consistent across implementations (ROC-AUC approximately 0.76-0.77). Prespecified pipelines achieved test ROC-AUC 0.770 (95% CI 0.767-0.773) nationally (R gradient boosting) and 0.762 (95% CI 0.746-0.777) in New York (Python AdaBoost). Soft-voting pipelines performed similarly (national 0.770; New York 0.757) and were treated as sensitivity benchmarks. Predicted probabilities were reasonably calibrated (Brier 0.118 nationally; 0.112 in New York), and higher scores among SCD-positive respondents persisted across survey years. Difficulty deciding, mental health, and functional health items ranked highest across permutation importance, SHAP, and GLMs. Respondents who reported no difficulty deciding (DECIDE = 2) had substantially lower odds of SCD than those who reported difficulty (DECIDE = 1; aOR approximately 0.13; FDR < 0.05). Training-only knowledge graphs likewise placed difficulty deciding nearest to SCD in both cohorts. Conclusions: A temporally locked, multi-pipeline BRFSS benchmark yields stable future-year SCD risk ranking, usable calibrated probability scores that remain separated by SCD status across survey years, and convergent interpretability signals. The open implementation supports reproducible surveillance-oriented machine learning for cognitive health.

20
Readability Assessment of Patient-Reported Measures Used During Heritable Cancer Genetic Testing

Adegbesan, A. C.; FitzGerald, L.; Dickinson, J. L.; Raspin, K.; Roydhouse, J.

2026-08-17 oncology 10.64898/2026.08.13.26360322 medRxiv
Top 0.5%
2.2%
Show abstract

Background: Patient-reported measures (PRMs), including patient-reported outcome and experience measures, capture patients perspectives on their health status and healthcare experiences. In cancer genetics, PRMs have been used to assess genetic knowledge, psychosocial outcomes, and decision-making. However, patients must understand these measures to provide useful information, an ability which is influenced by general and health literacy levels. Readability guidelines recommend that patient-facing materials be written at or below a Grade 6 level. This study evaluated the readability of PRMs used in a cancer genetic testing context. Objective: To assess whether PRMs used in heritable cancer genetic testing meet recommended readability levels using validated indices. Methods: PRMs were identified from a recent systematic review of PRMs used in heritable cancer genetic testing, which reported 83 instruments across eight categories. English-language PRMs containing structured question items and response scales were eligible for extraction and converted into plain text for analysis. Readability was assessed using four validated indices: Flesch Kincaid Grading Level (FKGL), FORd, CAylor, and STicht (FORCAST) formula, Flesch Reading Ease Score (FRES), and Simple Measure of Gobbledygook (SMOG) via an automated readability software. Descriptive analysis and numerical comparison evaluated readability levels across PRM categories and against the recommended Grade 6 reading level. Results: Sixty-five PRMs met the eligibility criteria, with most, including validated instruments, exceeding the recommended Grade 6 reading level. Across the eight categories, genetics-specific PRMs required the highest readability levels, indicating higher readability demands. Conclusions: Most PRMs, particularly those specific to genetics, do not meet readability guidelines. This may limit their accessibility to individuals with limited general and health literacy. Development of PRMs specific to genetics should consider strategies to improve readability, such as plain-language approaches and involvement of individuals with limited general or health literacy. Keywords: readability, patient-reported measures, cancer, genetic testing, health literacy